Papers with audio question answering benchmarks
AUDITA: A New Dataset to Audit Humans vs. AI Skill at Audio QA (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing audio question answering benchmarks emphasize sound event classification or caption-grounded queries. |
| Approach: | They propose a large-scale, real-world audio question answering benchmark to evaluate audio reasoning beyond surface-level acoustic recognition. |
| Outcome: | The proposed model achieves 32.13% accuracy while demonstrating comprehension of audio . state-of-the-art models perform poorly, with average accuracy below 8.86%. |